Philosophical Transactions of the Royal Society B
● The Royal Society
Preprints posted in the last 30 days, ranked by how well they match Philosophical Transactions of the Royal Society B's content profile, based on 51 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
Liang, Y.; Zuo, B.; Sun, Y.-B.
Show abstract
The Mutational Hazard Hypothesis (MHH) predicts that reduced effective population size weakens purifying selection and promotes transposable element (TE) accumulation. Because effective population size is difficult to estimate across broad taxonomic scales, body mass, generation time, and dN/dS are often used as proxies. We analyzed TE landscapes across 167 sauropsid genomes to test whether these proxies predict genomic TE proportion consistently across phylogenetic scales. All three showed strong scale dependence. Body mass was positively associated with TE proportion across clades, but this relationship broke down within clades: only snakes retained a significant negative Pearson correlation after multiple-testing correction, whereas the corresponding phylogenetically corrected slopes were not significant. Generation time showed a strong pooled association that disappeared within every major clade in both Pearson and PGLS analyses. The pooled dN/dS-TE relationship also disappeared after accounting for body mass and generation time in a structural equation model, with a robust within-clade association retained only in turtles. Lineage-specific TE dynamics, including LTR expansion in sea snakes, were not captured by these proxies. These results show that commonly used MHH proxies mainly reflect clade-level structure rather than consistent within-lineage mechanisms.
Yan, L.; Hu, S.; Ding, Y.; Jin, M.; Krasotkina, A.; Ren, L.; Liu, S.; Xiao, N. G.
Show abstract
Integrating auditory and visual cues is a hallmark of human speech perception, yet adults from East Asian backgrounds show less reliance on visual speech than their Western counterparts. The origins of this cultural difference, however, remain unknown. To investigate whether this divergence is established early in infancy, we examined audiovisual integration in 6- to 12- month-old White Canadian (n=111) and Chinese (n=115) infants using a novel paradigm measuring their perception of the McGurk effect. Across four experiments, we found a clear developmental divergence: Canadian infants showed a stable McGurk effect from 6 months onward, whereas Chinese infants showed a more protracted developmental trajectory, a cultural pattern that was further highlighted when their integration was challenged by other-race faces. These findings provide the first direct evidence that cultural differences in multisensory speech perception are established within the first year of life, suggesting that the brains strategy for binding sight and sound is shaped by early experience, with broad implications for theories of language acquisition and developmental science.
Shibasaki, S.
Show abstract
Rapid evolution allows populations to persist in environments where they would otherwise go extinct. This phenomenon, known as evolutionary rescue, is typically studied in the framework of biological evolution, yet adaptive traits can also arise and spread through cultural evolution. The present study developed a stochastic eco-evolutionary model to compare rescue probabilities through biological and cultural evolution. Transmission bias governed the rescue probability under cultural evolution by setting how readily a rare adaptive trait was copied. Conformity bias suppressed population persistence because a rare trait was the least likely to be copied. Content bias toward the adaptive trait enabled evolutionary rescue when social learning was rapid, but it typically yielded a lower rescue probability than biological evolution. Only anticonformity bias, together with a high social learning rate, exceeded the rescue probability of biological evolution by enabling the adaptive trait to be established more rapidly. These results demonstrate that transmission bias alters the demographic consequences of cultural evolution and highlight the importance of transmission processes in evolutionary rescue theory. Understanding how adaptive behaviours are socially transmitted may also improve predictions of animal population persistence and inform conservation efforts in rapidly changing environments.
Bentley, B. P.; Komoroske, L. M.; Santos, C. M.; Santos, A. J. B.; Argueta, E.; Quennessen, V.; Coppenrath, C. M.; Kynoch, C.; Saba, V. S.; Bellini, C.; Ventura, R. N. M. S.; White, J. W.; Fuentes, M. M. P. B.
Show abstract
Anthropogenic climate change is threatening global biodiversity, with sea turtles particularly vulnerable as offspring sex and developmental success are strongly influenced by incubation temperature. Behavioral plasticity, including the seasonal distribution of reproductive output, may provide short-term mechanisms for mitigating these impacts. Here, we investigated season-wide hatchling sex ratios and emergence success in a small population of green turtles (Chelonia mydas), tracking individual females across their nesting seasons. Sex ratios varied markedly through the nesting season, with later nests producing a greater proportion of male hatchlings. Moreover, sex ratios were relatively consistent among nests laid by individual females. Overall, females producing more nests over a season also produced more male offspring, suggesting that both nesting phenology and reproductive output influence individual contributions to future population demographics. Mechanistic models indicate that hatchling sex ratios have trended towards female-biased ratios (>80% female) over the past 50 years, and are projected to approach complete feminization by 2100 under continued warming. Although emergence success currently remains high (>85%), it is predicted to decline sharply after mid-century, with viable hatchling production falling to [~]30% by the end of the century. Models further show that maintaining contemporary sex ratios and emergence success will require unrealistically large delays in nesting phenology, and that even extreme shifts in phenology become ineffective by 2100. Together, these findings demonstrate that individual females can increase male hatchling production by nesting later and producing more nests, but behavioral plasticity alone is unlikely to offset the accelerating impacts of climate change on this population.
Spicher, L.; Huchard, E.; Lukas, D.
Show abstract
Classic socio-ecological theory predicts that males and females experience different sources and mechanisms of social competition. Whether these differences translate into sex-specific structural properties of dominance hierarchies remains unclear. Here, we compiled 156 dominance interaction matrices from 80 published studies and extracted three commonly used metrics - hierarchy steepness, linearity and the directional consistency index - to investigate the structural characteristics of male and female dominance hierarchies across primates. All three metrics were strongly affected by methodological and demographic variables. Steepness increased with the number of recorded interactions and group size, linearity decreased as matrices became sparser, and directional consistency declined with increasing numbers of interactions. Steepness covaried positively with both linearity and directional consistency, indicating that groups with steeper hierarchies also exhibited more linear and more directionally consistent relationships. We found no sex differences in steepness, linearity or directional consistency. These results suggest that current metrics primarily reflect variation in the sampling effort and the rate of interaction of the recorded behaviour and appear therefore not to capture potential sex differences in the forms of competition. Our findings highlight the need for alternative measures of power asymmetries that are less confounded by sampling effort and demographic variation to better understand how competition and conflict are structured across primate societies.
Sedibana, L.; Yessoufou, K.
Show abstract
Although cities are increasingly recognized as ecological islands, a unified framework explaining their susceptibility to alien plant invasion remains lacking. Using the most recent and comprehensive global dataset of urban alien plants, we modelled alien richness, mimicking island biogeography theory (IBT). Across all models, neither city size nor geographic isolation independently explained alien richness. Instead, richness was consistently associated with their interaction, supporting the central IBT prediction. However, the strength of this interaction depends on how city size was quantified, with socio-economic dimensions exhibiting stronger positive interactions with geographic isolation than physical measures of city size. Introduction-hub identity further modified these relationships. North America was the only hub for which the interaction between city size and isolation was consistently weakened, indicating that donor regions of alien plants are not ecologically equivalent. Simulations of simultaneous increases in city size and isolation showed that larger, more connected cities generally accumulated more alien plants despite increasing geographic distance, but the magnitude and direction of these responses are hub dependent. Our findings inspire an extension of classical IBT to a mechanistic explanation for global variation in urban alien plant richness in this increasingly urbanized and globally connected world.
Tamai, Y.; Matsumoto, J.; Löschner, J.; König, L.; Kaneko, T.; Inoue, K.-i.; Toda, K.; Hage, S.
Show abstract
Human conversation depends on the continuous integration of vocal exchanges with visual and spatial cues, yet the evolutionary origins of this multimodal coordination remain poorly understood. Although vocal turn-taking has been documented across many animal species, studies have largely examined vocal exchanges in isolation from the accompanying social dynamics. Using acoustic localization and 3D pose tracking in freely interacting marmoset pairs, we simultaneously quantified vocal behaviour and social interactions during natural communication. We found that vocal turn-taking is dependent on distinct multimodal behavioural states defined by head orientation, spatial proximity, and ongoing social interaction. While call features did not reliably predict turn-taking, these behavioural dynamics predicted whether vocal exchanges developed into turn-taking or terminated after isolated calls. Our findings reveal that primate vocal communication is fundamentally organized by multimodal behavioural coordination rather than by vocal signals alone, providing an evolutionary framework for understanding the origins of human conversation.
Dos Santos, M.; Ohtsuki, H.; Mullon, C.
Show abstract
Reputation plays a major role in supporting cooperation among unrelated individuals through indirect reciprocity. By helping others, individuals build a good personal reputation and receive greater benefits from future partners. Most models of indirect reciprocity assume that a person's reputation reflects only their own behaviour. Yet in many societies, people are also judged by their family's reputation. How family reputation affects the evolution of cooperation, and whether reliance on it can itself evolve, remain unclear. Here we show that reputation inheritance expands the conditions under which indirect reciprocity favours cooperation, increasing helping and favouring greater reciprocity. Greater reciprocity in turn favours stronger reliance on inherited reputation, creating a positive feedback that stabilises cooperation, especially when interactions are infrequent or personal behaviour is difficult to observe. This feedback arises because cooperation generates future benefits both for the individual, through their personal reputation, and for their descendants, through inherited reputation. Reputation inheritance thereby provides a route via which kin selection and reciprocity, often treated as alternative explanations for cooperation, can reinforce one another. Our model helps explain why family-based reputation occurs across diverse human societies and provides an evolutionary framework for studying phenomena organised around family standing, including kin-based institutions, feuds between families and honour-based violence within them.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Chakraborty, A.; Agashe, D.
Show abstract
Evolution of antibiotic resistance is a major global public health problem. Rapid emergence of antibiotic resistance is often linked to hypermutator bacteria with defective DNA repair, leading to high mutation rates that are broadly advantageous. However, depending on which DNA repair pathway is dysfunctional, mutators may sample only specific types of mutations at a higher rate. Thus, their mutation spectrum can be biased towards specific mutation types, influencing the identity of resistance mutations. Under strong antibiotic selection, an overall high mutation rate should generally shorten the time to sample a resistance mutation and increase the probability of resistance. However, recent work suggests that the mutation rate for specific types of mutations in target genes that drive high resistance is more important than the overall mutation rate. To systematically test this prediction, we exposed Escherichia coli mutators with varying mutation rates and spectra to antibiotics targeting different cellular functions. For each strain, we determined the highest antibiotic concentration at which resistance could emerge overnight, quantifying both the magnitude and probability of resistance. High level antibiotic resistance was generally better predicted by specific rather than overall mutation rate, and resistance mutations matched the mutation spectrum of the respective mutator. Despite the varying magnitude of resistance, at the highest antibiotic concentration survived by each strain, the respective resistance mutations were generally costly in the absence of antibiotic. Given that mutators often arise in laboratory, natural, and clinical settings under antibiotic selection, we suggest that their mutation spectra deserve more attention.
Pringle, J. M.; Lush, W. G.; Byers, J. E.
Show abstract
After introduction, many non-native marine species are dispersed planktonically. Secondary spread within the non-native range has been shown to prevent the establishment of the introduced species if the advection of larvae prevents sufficient return of larvae to maintain the population in the face of competition with native species. However, those studies have largely neglected the effects of spatial variation in alongshore larval transport. We examine the introduction of a novel species with planktonic dispersal into a more realistic coastal environment which includes spatial variation in larval transport estimated from the Mercator Ocean 1/12th degree global circulation model. The introduction may either be from a distant habitat, or through range expansion. We find that there are locations in the global coastal ocean where introduced species are more likely to persist because of spatial variation of coastal currents. These include regions where alongshore larval transport diverges, such as estuaries. The location where a non-native species is introduced may not be where it flourishes - it cannot be assumed that the region where invading species are first noticed to be abundant is the region where it was introduced. We extend closed-population theory to open coastal systems to estimate persistence as a function of local circulation, habitat extent, and the competitive advantage of the introduced species. Software is provided which allows the estimations of regions where introduced species are more likely to persist and flourish as a function of larval depth behavior, planktonic duration and release timing.
Demirel, B.; Parr, T.; Saleh, Y.; Jackson, E. S.; Denison, T.; Manohar, S. G.
Show abstract
Adults who stutter can speak fluently when speech is not addressed to another person, but stuttering emerges when they aim to convey information to a listener. The value of the information being conveyed to the listener also affects the likelihood of stuttering. Why should the mere absence of a listener neutralise a profound motor deficit, and why does a word's predictability affect whether it is spoken fluently? To resolve this socio-motor paradox, we develop a computational model of stuttering within an active inference architecture. The model represents the communicative context, including whether a listener is present and whether the agent is speaking or listening. It was designed around two candidate mechanisms for stuttering, a prior for silence and rigid phoneme sequencing precision. Using both, the model produced fluent private speech and more stuttering-like events during social speech. In the same parameter regime, the model also showed more stuttering-like events on words with higher information value, and produced a word-length effect, in which disfluency increased with longer words. To our knowledge, this is the first model of stuttering to generate both the private speech and the information-value effect from inferred communicative context. By representing the listener as a hidden state that makes the sensory consequences of resuming speech ambiguous, the model offers a computational link between social cognition and speech-motor instability, and suggests that speech fluency depends on whether the speaker believes anyone is present. Clinically, it may offer testable hypotheses and a route to personalising treatment, since the same overt severity can arise from different combinations of parameters.
Lin, H.-w.; Hernandez, C.; Jaggi, H.; ZUO, W.; Tuljapurkar, S. D.; Salguero-Gomez, R.
Show abstract
The performance of any natural population in variable environments depends on contemporaneous changes in its vital rates (e.g., survival, reproduction) as well as legacies carried by its population structure. Yet whether the relative contribution of these two pathways can be predicted from life history remains unknown. Here, we use stochastic simulations of 1,986 matrix population models from 137 species to quantify the contribution of transient dynamics to variation in population growth rate, and test its associations with key life history traits. Longer generation times were associated with reductions in transient contributions, contrary to theoretical expectations. Greater stage-specific survival heterogeneities were associated with increases in transient contributions, whereas greater iteroparity was associated with decreases in plants but increases in animals. These associations were robust to body size, phylogenetic relationships, and vital-rate variability. Life history traits therefore provide a strong predictor for when population structure shapes population responses to environmental variability.
Mwangi, B.; Wu, M.-J.; Mansour, R.; Anzueto, G.; Pagan, A. F.
Show abstract
Background Naturalistic audiovisual recordings of caregiver-child interactions contain rich developmental signals. However, extracting interpretable clinical measures requires resource-intensive manual coding. To address this bottleneck, we evaluated natural-language queries for retrieving specific behavioral moments from these recordings, applying multimodal embeddings as an automated evidence-selection layer. Methods We compared three embedding models (Jina Embeddings v5 Omni, LanguageBind, and Wave7B) for natural-language retrieval directly from audio and video streams, bypassing transcript text. We assessed performance across 27 behavioral targets in 277 caregiver-child recordings (14, 24, and 36 months of age) from the Early Head Start Talkbank corpus, yielding 7,479 recording-target queries. Results Jina Embeddings v5 Omni achieved the highest top-10 retrieval success (text-to-audio 38.3%; text-to-video 36.4%), ahead of LanguageBind (37.0%; 34.5%) and Wave7B (36.1%; 35.0%). Across models, retrieval was substantially more successful for common targets than for rare vocal and gestural behaviors, such as pointing and babbling. By analyzing the spoken words within the retrieved audio clips, we found that Jina accurately ranked the children by their relative vocabulary size at each age (Spearman = 0.68, 0.82, and 0.90 at 14, 24, and 36 months). However, the model severely underestimated the total number of unique words each child used throughout the full session. Conclusion Multimodal embeddings can successfully pinpoint important developmental behaviors and speech patterns within lengthy caregiver-child recordings. However, these systems still struggle to locate rare events. Additionally, while they can accurately rank children by relative vocabulary size, they fail to measure a child's complete vocabulary. We conclude that these models are currently best suited for automated evidence-selection to prioritize relevant segments for expert interpretation rather than acting as an independent replacement for manual behavioral coding or language assessment. Improving the detection of infrequent behaviors and validating these models across external datasets are essential next steps before real-world clinical deployment.
Kijima, A.; Okumura, M.; Shima, H.; Kallen, R. W.; Richardson, M. J.; Yamamoto, Y.
Show abstract
Predicting patterns of behavioural coordination that emerge in small interacting groups is challenging because goal-directed social action is shaped by complex reciprocal and compensatory dynamics. In this study, we examined whether formal symmetry principles derived from group theory could explain coordination patterns among children performing a triadic jumping task. We investigated how geometric symmetries of the task environment and dispositional (a)symmetries associated with leader-follower tendencies jointly constrain collective behaviour. Forty-seven children were classified into symmetric or asymmetric triads based on teacher evaluations of leadership dispositions. Each triad completed multiple trials of a synchronized jumping game requiring movement between adjacent hoops arranged in triangular or square configurations. Results showed that temporal asymmetries in inter-child movement (first, second, or last to jump) were consistent with group-theoretic predictions. In triangular configurations, observed asymmetries corresponded to the highest-order subgroup defined by task and dispositional symmetries. These findings demonstrate that environmental symmetry exerts a hierarchically dominant constraint on collective coordination, within which actor dispositional (a)symmetries further modulate emerging patterns.
Aggarwal, K.; Samad, I.; Thaker, M.; Shanker, K.
Show abstract
Mixed-species groups (MSGs) pose a particular challenge for our understanding of sociality in animals. Though MSGs are widespread social assemblages that form to enhance foraging success and reduce predation risk of participants, the role of traits in mediating grouping has received less attention. In particular, the role of body colour has not been tested quantitatively, despite the fact that visual similarity can reduce individual predation risk. Here, we examine whether plumage colour structures mixed-species bird flocks (MSFs) at a global scale. Using data spanning four continents, we developed a new metric that quantifies colour similarity among flock participants and compared observed flocks to null assemblages constructed from all flocking species at each site. We further examined whether MSF participants represented a colour subset of the available colours in the regional species pool. We found striking evidence that birds in MSFs were more similar in colour than expected by chance across all sites, indicating that plumage colour is a non-random structuring trait that shapes assembly of flocks globally. The strength and prevalence of colour structuring varied across geographies, but not flock size. Within communities, MSF participants differed systematically in colour composition from the regional species pool, occupying a restricted region of colour space dominated by yellow and brown plumage. Thus, plumage colour affects MSFs influencing both overall flock participation as well as species co-occurrence within flocks. Our findings illustrate the importance of visual traits in structuring interspecific social systems, by highlighting that birds of a feather do indeed flock together.
Gorenshtein, A.; Jia, E. L.; Omar, M.; Brook, O. R.; Ahmed, M.; Kruskel, J. B.; Barash, Y.; Klang, E.
Show abstract
Safety alignment should persist while a language model performs a task. We tested whether a single-patient triage task suppressed a warning about a second patient. Each case centered on Patient 1; Patient 2's urgent problem appeared only in passing. Sixteen models saw each case twice: once as a general assistant and once while producing a triage record for Patient 1. As general assistants, models warned the caller in 87% of cases; under the task, they did so in 21%. Every model showed a significant decrease. Yet under the task, the record still mentioned Patient 2 in 76% of cases and recommended urgent care in 67%. Across 15 open-weight models, repeating the emergency-care instruction raised the warning rate only to 29%; moving the message-to-caller field to the top raised it to 36%. Current safety alignment did not reliably persist under task assignment.
Setiono, F. J.; Ho, E.; Lambert, W. M.
Show abstract
Effective mentorship is essential for strengthening the STEMM (Science, Technology, Engineering, Mathematics, and Medicine) workforce, yet empirical evidence on how mentorship networks are structured and linked to career success remains limited. Here, we analyze mentorship networks among recipients of NIH career development (K) awards to characterize network size, mentor roles, and their associations with mentee-reported outcomes, including potential variation by sociodemographic characteristics. We found that K-awardees rely on mentors beyond their primary advisor, who play varying roles beyond being a Research mentor. Different mentor roles led to different types of mentoring outcomes; while Research mentors were associated with research-related outcomes such as Publications and Grants, career- and psychosocial-related mentoring outcomes were more likely to come from other types of mentors, such as Coaches, Connectors, and Sponsors. Larger networks, as well as having Peer and Identity mentors are additively beneficial for researchers who identify as underrepresented in science more than their counterparts. This study provides large-scale evidence on how mentorship network configurations relate to early-career grant success.
Rony, A. R.; Nahin, K. S. A.; Islam, T.; Asha, A. S.; Hossen, A.
Show abstract
Caesarean section in Bangladesh reached 51.8% of deliveries in 2025, and elective caesarean, meaning caesarean before labour began, reached 31.6%. Risk models built on national household surveys are increasingly proposed for pointing audit toward places where scheduled surgery is outrunning clinical need, but they are usually validated in ways that flatter them. Using the 2025 Bangladesh Multiple Indicator Cluster Survey, we developed four models on 9,538 women (logistic regression, elastic net, random forest, gradient boosting) and ran the same procedure under three validation designs: random five-fold cross-validation; five-fold cross-validation grouped by sampling cluster; and leave-one-division-out cross-validation. We also tested transfer between the 2019 and 2025 rounds and audited subgroup calibration. No model improved on logistic regression by a margin worth acting on: the area under the receiver operating characteristic curve ranged from 0.724 to 0.736 under cluster-grouped validation, a spread of 0.012. Validation design mattered far more than the algorithm. Grouping folds by sampling cluster changed discrimination by at most 0.0004, this survey contributing a median of 3 eligible women per enumeration area. Withholding a whole division cost 0.044 to 0.060, more than 100 times as much, and still cost 0.033 to 0.056 after the strongest predictor, an outcome-derived district rate, was removed from every model. A model fitted to 2019 data lost 0.083 when applied to 2025, and the two rounds agreed only moderately on which predictors mattered (Spearman rank correlation 0.61). Calibration held in every wealth quintile, both residence categories and seven of eight divisions; Sylhet was the exception. Elective caesarean is predictable from routine survey items, but that predictability is local. Cross-validation, including cluster-aware cross-validation, does not measure what a model would do in a district it has never seen; a geographic holdout is the cheapest design that does.
Mandich, A.; Koirala, S.; Westen, S.; Adhikari, S.; Acharya, A.; Shrestha, A.
Show abstract
Language discordance can impede community-based research and health communication where trained interpreters are limited. Although multimodal artificial intelligence systems can provide real-time spoken translation, performance with under-resourced languages during spontaneous field interactions remains poorly characterized. We evaluated ChatGPT-4o during bidirectional English-Nepali voice translation in a community setting near Dhulikhel Hospital, Nepal. In this cross-sectional field study, 30 primarily Nepali-speaking adults were recruited by convenience sampling. ChatGPT-4o mediated conversations using standardized English questions and spontaneous Nepali responses. A bilingual Nepali-English reviewer assessed 485 translated utterances using a 3-point accuracy scale and an inductively developed framework for translation and conversational deviations. Of 485 translations, 282 (58.1%) received the highest accuracy rating, 134 (27.6%) a moderate rating, and 69 (14.2%) the lowest. Mean accuracy was higher for English-to-Nepali than Nepali-to-English translation (2.63 {+/-} 0.53 vs 2.23 {+/-} 0.86); 63 of 69 low-accuracy translations (91.3%) occurred in the Nepali-to-English direction. Among 329 deviation tags, the most frequent were distortion of intended meaning (17.1%), overly formal or unnatural phrasing (14.7%), omission (14.2%), and addition of content (11.5%). Some fluent outputs substantially altered meaning or introduced information not expressed by the speaker. ChatGPT-4o demonstrated potential for real-time English-Nepali communication but also produced errors that could alter interpretation of participant responses. Accuracy was lower and more variable for Nepali-to-English translation; however, translation direction was confounded with input type because Nepali inputs were spontaneous and English inputs standardized, limiting conclusions about directional performance. These findings support cautious use for low-stakes conversational exchange and human verification when errors could affect research validity, clinical decisions, or participant understanding. As multimodal AI evolves, performance should be reevaluated across languages, real-world conditions, and model versions, with bilingual oversight and community partnership remaining central to responsible use.